Papers with multi-modal models
Learning the Effects of Physical Actions in a Multi-modal Environment (2023.findings-eacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are trained on large corpora of disembodied texts. |
| Approach: | They propose a multi-modal task of predicting the outcomes of actions solely from realistic sensory inputs (images and text). They extend an LLM to model latent representations of objects to better predict action outcomes in an environment. |
| Outcome: | The proposed model can capture commonsense when augmented with visual information and generalize and learn commonsensical reasoning better. |
Visio-Linguistic Brain Encoding (2022.coling-1)
Copied to clipboard
| Challenge: | Existing studies have failed to explore co-attentive multi-modal modeling for visual and text reasoning. |
| Approach: | They propose to use image and multi-modal Transformers to reconstruct fMRI brain activity . they use two popular datasets to study visual and text reasoning . |
| Outcome: | The proposed model outperforms existing models on two popular datasets . the results raise the question whether visual processing is affected implicitly by linguistic processing . |
Query-aware Multi-modal based Ranking Relevance in Video Search (2023.emnlp-industry)
Copied to clipboard
| Challenge: | Existing relevance ranking methods focus on text modality, incapable of fully exploiting cross-modal cues present in video. |
| Approach: | They propose a QUery-Aware pre-training model with multi-modality that integrates video tag information as alignment targets and enhances ranking optimization method based on ordinal regression. |
| Outcome: | The proposed model significantly improves video search performance. |
Abstract Visual Reasoning with Tangram Shapes (2022.emnlp-main)
Copied to clipboard
| Challenge: | We use tangrams as stimuli in cognitive science to study abstract visual reasoning . pre-trained weights demonstrate limited abstract reasoning, we observe . |
| Approach: | They propose a resource for studying abstract visual reasoning in humans and machines . they use tangram puzzles as stimuli to create an annotated dataset with >1k distinct stimuli . |
| Outcome: | The proposed resource is visually and linguistically richer than previous resources . pre-trained weights demonstrate limited abstract reasoning, the authors note . |
Visual Zero-Shot E-Commerce Product Attribute Value Extraction (2025.naacl-industry)
Copied to clipboard
| Challenge: | Existing zero-shot product attribute value extraction approaches require sellers to manually provide product descriptions. |
| Approach: | They propose a cross-modal zero-shot attribute value generation framework based on CLIP that uses product images as inputs for zero- shot inference. |
| Outcome: | The proposed framework significantly outperforms other vision-language models for zero-shot attribute value extraction. |
GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets are too challenging for direct model learning or suffer from misalignment between text and images. |
| Approach: | They propose a pipeline that leverages GPT-4 and GPT4V to generate geometry problems with aligned text and images, facilitating model learning. |
| Outcome: | The proposed pipeline generates 4.9K geometry problems with aligned text and images, facilitating model learning. |
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) and multi-modal models (MMs) have demonstrated remarkable capabilities in problem-solving, but their proficiency in tackling geometry math problems has not been thoroughly evaluated. |
| Approach: | They propose a benchmark to evaluate the performance of large language models and multi-modal models in solving geometry math problems. |
| Outcome: | The proposed model achieves 55.67% accuracy on main subset but only 6.00% accuracy on hard subset. |
TurkingBench: A Challenge Benchmark for Web Agents (2025.naacl-long)
Copied to clipboard
Kevin Xu, Yeganeh Kordi, Tanay Nayak, Adi Asija, Yizhong Wang, Kate Sanders, Adam Byerly, Jingyu Zhang, Benjamin Van Durme, Daniel Khashabi
| Challenge: | TurkingBench is a benchmark consisting of tasks presented as web pages with textual instructions and multi-modal contexts. |
| Approach: | They propose to use HTML pages to perform various annotation tasks on crowdsourcing platforms. |
| Outcome: | The proposed model outperforms other models on the TurkingBench benchmark. |
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Despite the impressive capabilities of large multi-modal models, their effectiveness in handling complex tasks has been limited by the prevailing singlestep reasoning paradigm. |
| Approach: | They propose a visuallygrounded object-centric Chain-of-Thought reasoning framework for LMMs that is based on a multi-modal interleaved and aligned representation of object concepts. |
| Outcome: | The proposed model outperforms SOTA models in CLEVR and EmbSpatial benchmarks. |
Selectively Answering Visual Questions (2024.findings-acl)
Copied to clipboard
| Challenge: | Large multi-modal models (LMMs) are capable of visual question answering (VQA) with unprecedented accuracy. |
| Approach: | They propose a calibration score that can be used to quantify uncertainty in visual question answering models. |
| Outcome: | The proposed calibration score is better calibrated than in text-only models for in-context learning. |
Uncovering Visual-Semantic Psycholinguistic Properties from the Distributional Structure of Text Embedding Space (2025.acl-long)
Copied to clipboard
| Challenge: | Imageability and concreteness are psycholinguistic properties that link visual and semantic spaces. |
| Approach: | They propose an unsupervised measure that quantifies sharpness of peaks in an image-caption dataset. |
| Outcome: | The proposed method is more robust than existing methods and predicts these properties for classification. |
Is the Red Square Big? MALeViC: Modeling Adjectives Leveraging Visual Contexts (D19-1)
Copied to clipboard
| Challenge: | gradable adjectives of size are relative, i.e., determined by the context. |
| Approach: | They propose to model how the meaning of gradable adjectives of size can be learned from visually-grounded contexts by using four tasks to determine whether an object is ‘big’ or ‘small’. |
| Outcome: | The proposed model can learn subtending the meaning of size adjectives, but their performance decreases while moving from simple to more complex tasks. |
UICoder: Finetuning Large Language Models to Generate User Interface Code through Automated Feedback (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing approaches to improve UI code generation rely on expensive human feedback or distilling a proprietary model. |
| Approach: | They propose to use automated feedback to guide large language models to generate UI code . they use a large synthetic dataset to generate improved models and refine them . |
| Outcome: | The proposed model outperforms baseline models and larger proprietary models . the model outpersforms models with automated metrics and human preferences . |
Aligning to What? Limits to RLHF Based Alignment (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing studies on RLHF and covert and overt biases in large language models are unclear . et al. analyzed off-the-shelf language models to evaluate their overt and cover racial biase . |
| Approach: | They evaluate the relationship between reinforcement learning from human feedback and biases in large language models. |
| Outcome: | The proposed approach can be used to mitigat covert biases, the authors show . they found that the RLHF approach calcifies model biase . |
Can Textual Unlearning Solve Cross-Modality Safety Alignment? (2024.findings-emnlp)
Copied to clipboard
Trishna Chakraborty, Erfan Shayegani, Zikui Cai, Nael Abu-Ghazaleh, M. Salman Asif, Yue Dong, Amit Roy-Chowdhury, Chengyu Song
| Challenge: | integrating new modalities into large language models creates new attack surface . existing safety training techniques like SFT and RLHF are not feasible in multi-modal settings . |
| Approach: | They explore whether unlearning in the textual domain can be effective for cross-modality safety alignment. |
| Outcome: | The proposed approach reduces the Attack Success Rate (ASR) to less than 8% and preserves the utility. |
MindBridge: Scalable and Cross-Model Knowledge Editing via Memory-Augmented Modality (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing knowledge editing methods overfit to specific models, causing edited knowledge to be discarded during each LLM update and requiring frequent re-editing. |
| Approach: | They propose a solution that allows editors to edit knowledge in multiple LLMs at the same time. |
| Outcome: | The proposed solution performs better even in editing tens of thousands of knowledge entries and can adapt to different LLMs. |
Do Current Video LLMs Have Strong OCR Abilities? A Preliminary Study (2025.coling-main)
Copied to clipboard
| Challenge: | a new benchmark evaluates video-based optical character recognition (Video OCR) performance of multi-modal models in videos . the benchmark aims to improve video LLMs' ability to extract text from video content . previous benchmarks have focused on video QA, but not video-related QA. |
| Approach: | They propose to evaluate the video OCR performance of multi-modal models in videos . they use a semi-automated approach that integrates the OCR ability of image LLMs with manual refinement . |
| Outcome: | The proposed benchmark includes 1,028 videos and 2,961 question-answer pairs . it integrates the OCR ability of image LLMs with manual refinement . |
VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape Rooms (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on embodied agents have addressed the importance of exploration in environments where tasks and solutions are not predefined. |
| Approach: | They propose a virtual escape room that evaluates AI models in a dynamic environment . they propose to integrate memory management and reasoning into the simulation . |
| Outcome: | The proposed model improves in dynamic and exploration-driven environments by integrating memory management and reasoning. |
DetGPT: Detect What You Need via Reasoning (2023.emnlp-main)
Copied to clipboard
Renjie Pi, Jiahui Gao, Shizhe Diao, Rui Pan, Hanze Dong, Jipeng Zhang, Lewei Yao, Jianhua Han, Hang Xu, Lingpeng Kong, Tong Zhang
| Challenge: | Recent advances in the field of computer vision have enabled more effective and sophisticated interactions between humans and machines. |
| Approach: | They propose a reasoning-based object detection paradigm that leverages state-of-the-art multi-modal models and open-vocabulary object detectors to perform reasoning within the context of the user’s instructions and the visual scene. |
| Outcome: | The proposed method enables users to interact with the system using natural language instructions, allowing for a higher level of interactivity. |
BloomVQA: Assessing Hierarchical Multi-modal Comprehension (2024.findings-acl)
Copied to clipboard
Yunye Gong, Robik Shrestha, Jared Claypoole, Michael Cogswell, Arijit Ray, Christopher Kanan, Ajay Divakaran
| Challenge: | Recent advances of machine intelligence solutions have demonstrated tremendous success in a wide range of language and multi-modal tasks over diverse domains. |
| Approach: | They propose a VQA dataset to facilitate comprehensive evaluation of large vision-language models on comprehension tasks. |
| Outcome: | The proposed dataset shows improved accuracy over all comprehension levels and a tendency to bypass visual inputs especially for higher-level tasks. |
Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? (2023.emnlp-main)
Copied to clipboard
| Challenge: | Pre-trained vision and language models have demonstrated state-of-the-art capabilities over existing tasks involving images and texts. |
| Approach: | They analyze a visual question answering dataset tailored for info-seeking questions . they show that pre-trained visual and language models can use fine-grained knowledge . |
| Outcome: | The proposed dataset elicits models to use fine-grained knowledge learned during pre-training. |
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows (2025.emnlp-main)
Copied to clipboard
Kirolos Ataallah, Eslam Mohamed Bakr, Mahmoud Ahmed, Chenhui Gou, Khushbu Pahwa, Jian Ding, Mohamed Elhoseiny
| Challenge: | Existing benchmarks fail to test the full range of cognitive skills needed to process long-form videos . |
| Approach: | They propose a benchmark to evaluate models' ability to process long-form videos rigorously. |
| Outcome: | The benchmark measures the cognitive skills of models in understanding long-form videos . it offers the largest set of question-answer pairs for long video comprehension . |
Cross-MoE: An Efficient Temporal Prediction Framework Integrating Textual Modality (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing models ignore dynamic and different relations between time series patterns and textual features, which leads to poor performance in temporal-textual feature fusion. |
| Approach: | They propose a temporal-textual fusion framework that replaces Cross Attention with Cross-Ranker to reduce computational complexity and enhances modality-aware correlation memorization with Mixture-of-Experts (MoE) networks to tolerate the distributional shifts in time series. |
| Outcome: | The proposed framework reduces MSE by 8.78% compared to the current SOTA model and requires only 75% of computational overhead and 12.5% of activated parameters. |
MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models (2026.acl-long)
Copied to clipboard
Yang Shi, Yifeng Xie, Minzhe Guo, Liangsi Lu, Mingxuan Huang, Jingchao Wang, Zhihong Zhu, Boyan Xu, Zhiqi Huang
| Challenge: | Recent advances in vision-language models have improved performance in multi-modal learning. |
| Approach: | They propose a multi-modal benchmark that embeds a single coherent reasoning error in 1997 samples. |
| Outcome: | The proposed benchmark is based on a set of 1997 samples embedding a single coherent reasoning error. |